operating system
The first 5 Googlebook laptops are here, with an OS that blends Android and Gemini
Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Look Up Versus Say More Mashable Selects Creator Playbook In My Bag Trending Now Back to School Good Connection: Uplifting stories for a digital age Switch Off Mashable Voices All Series Say that five times fast. Alex Perry is a tech reporter at Mashable who primarily covers video games and consumer tech. Alex has spent most of the last decade reviewing games, smartphones, headphones, and laptops, and he doesn't plan on stopping anytime soon. He is also a Pisces, a cat lover, and a Kansas City sports fan. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals.
Circle CEO Jeremy Allaire on Arc, the Future of Quantum and the Agentic Economy
Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens.
AtAtT!T" O!O" Al-to-AlE!E" E# E$ AtT!T" O!O"FFNAtGaT!T" O!O" GaGaE!E" E# E$ Al-to-Al(a(b(clllltetetete))) tetete((DCnnnnDTMoiomttttrsiiiiiaoooopsnnnnansbEttricnihbfe)o)urtmede rMoE
The computational sparsity of Mixture-of-Experts (MoE) models enables sublinear growth in compute cost as model size increases, thus offering a scalable path to training massive neural networks. However, existing implementations suffer from low GPU utilization, significant latency overhead, and a fundamental inability to leverage task locality, primarily due to CPU-managed scheduling, host-initiated communication, and frequent kernel launches. To overcome these limitations, we develop FlashMoE, a fully GPU-resident MoE operator that fuses expert computation and inter-GPU communication into a single persistent GPU kernel. FlashMoE enables fine-grained pipelining of dispatch, compute, and combine phases, eliminating launch overheads and reducing idle gaps. Unlike existing work, FlashMoE obviates bulk-synchronous collectives for one-sided, device-initiated, inter-GPU (R)DMA transfers, thus unlocking payload efficiency, where we eliminate bloated or redundant network payloads in sparsely activated layers. When evaluated on an 8-H100 GPU node with MoE models having up to 128 experts and 16K token sequences, FlashMoE achieves up to 9 higher GPU utilization, 6 lower latency, 5.7 higher throughput, and 4 better overlap efficiency compared to state-of-the-art baselines--despite using FP32 while baselines use FP16. FlashMoE shows that principled GPU kernel-hardware co-design is key to unlocking the performance ceiling of large-scale distributed ML.
Swedish Death Cleaning, but for Your Digital Life
The art of ordering and culling your possessions before you die should extend to your documents, photos, and digital accounts. Digital generated image of semi transparent multiple data server discs on white background. After Adam Liljenberg's grandmother died, his grandfather was ready to downsize and move into an assisted living facility. As Swedes, they were familiar with Swedish death cleaning, the idea that as you near the end of life, you declutter and organize your belongings so as not to burden those who survive you. When Liljenberg arrived to help his grandfather sort through his possessions, he didn't expect to be rescuing digital photos off a phone full of malware.
All Windows 11 PCs Will Get These Advanced Copilot AI Features
As Windows 10 Support Ends, Microsoft Is'Rewriting' Windows 11 Around AI All Windows 11 users will soon be able to talk to the Copilot AI assistant more easily via voice, and Copilot Vision can understand the context of your screen. Microsoft saved its most powerful AI tools for paying customers in the first phase of its AI evolution. Now, the company has announced a series of Copilot features coming to all Windows 11 PCs, including Voice, Copilot Vision, and Copilot Actions. Alongside the update, Microsoft is launching an ad campaign to expose people to these new features. Windows 10 support ended on October 14, and we're about to see a wave of people upgrade to Windows 11; Microsoft seems intent on putting advanced Copilot features at the fingertips of as many people as possible--and convincing them they're worth using.
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
Bitchebe, Stella, Balmau, Oana
A large body of research has employed Machine Learning (ML) models to develop learned operating systems (OSes) and kernels. The latter dynamically adapts to the job load and dynamically adjusts resources (CPU, IO, memory, network bandwidth) allocation to respond to the actual user demand. What this work has in common is that it utilizes ML to improve kernel decisions. To this day, and to the best of our knowledge, no work has taken the opposite direction, i.e., using OS to improve ML. While some work proposes applying system-level optimizations to ML algorithms, they do not tailor the OS to adapt to the ML context. To address this limitation, we take an orthogonal approach in this paper by leveraging the OS to enhance the performance of ML models and algorithms. We explore the path towards an ML-specialized OS, MaLV-OS. MaLV-OS rethinks the OS architecture to make it specifically tailored to ML workloads, especially in virtualized clouds, which are now widely used to run ML applications. MaLV-OS envisioned architecture includes (1) a micro-kernel, Micro-LAKE, which allows kernel space applications to use the GPU, and (2) an MLaaS (ML as a Service) subsystem that gathers ML models to help Micro-LAKE with memory management and CPU scheduling. MaLV-OS architecture also offloads system-sensitive parts of the models to the OS, to lighten the model complexity and programming, and speed up its execution. Finally, MaLV-OS integrates an open-source GPU virtualization software, merged directly into the hypervisor. For more flexibility, MaLV-OS vision is to enable the virtual machine to dynamically select MLaaS policies that can improve the performance of the model the user is running. Because MLaaS is designed as loadable kernel modules, the MaLV-OS architecture enables the dynamic addition of new capabilities to the MLaaS subsystem.
Composable OS Kernel Architectures for Autonomous Intelligence
Singh, Rajpreet, Kothari, Vidhi
As intelligent systems permeate edge devices, cloud infrastructure, and embedded real-time environments, this research proposes a new OS kernel architecture for intelligent systems, transforming kernels from static resource managers to adaptive, AI-integrated platforms. Key contributions include: (1) treating Loadable Kernel Modules (LKMs) as AI-oriented computation units for fast sensory and cognitive processing in kernel space; (2) expanding the Linux kernel into an AI-native environment with built-in deep learning inference, floating-point acceleration, and real-time adaptive scheduling for efficient ML workloads; and (3) introducing a Neurosymbolic kernel design leveraging Category Theory and Homotopy Type Theory to unify symbolic reasoning and differentiable logic within OS internals. Together, these approaches enable operating systems to proactively anticipate and adapt to the cognitive needs of autonomous intelligent applications.
LithOS: An Operating System for Efficient Machine Learning on GPUs
Coppock, Patrick H., Zhang, Brian, Solomon, Eliot H., Kypriotis, Vasilis, Yang, Leon, Sharma, Bikash, Schatzberg, Dan, Mowry, Todd C., Skarlatos, Dimitrios
The surging demand for GPUs in datacenters for machine learning (ML) has made efficient GPU utilization crucial. However, meeting the diverse needs of ML models while optimizing resource usage is challenging. To enable transparent, fine-grained GPU management that maximizes utilization and energy efficiency while maintaining strong isolation, an operating system (OS) approach is needed. This paper introduces LithOS, a first step toward a GPU OS. LithOS includes the following new abstractions and mechanisms for efficient GPU resource management: (i) a novel TPC Scheduler that supports spatial scheduling at the granularity of individual TPCs, unlocking efficient TPC stealing between workloads; (ii) transparent kernel atomization to reduce head-of-line blocking and enable dynamic resource reallocation mid-execution; (iii) a lightweight hardware right-sizing mechanism that determines the minimal TPC resources needed per atom; and (iv) a transparent power management mechanism that reduces power consumption based on in-flight work behavior. We implement LithOS in Rust and evaluate its performance across extensive ML environments, comparing it to state-of-the-art solutions from NVIDIA and prior research. For inference stacking, LithOS reduces tail latencies by 13x compared to MPS; compared to the best SotA, it reduces tail latencies by 3x while improving aggregate throughput by 1.6x. In hybrid inference-training stacking, LithOS reduces tail latencies by 4.7x compared to MPS; compared to the best SotA, it reduces tail latencies 1.18x while improving aggregate throughput by 1.35x. Finally, for a modest performance hit under 4%, LithOS's right-sizing provides a quarter of GPU capacity savings on average, while for a 7% hit, its power management yields a quarter of a GPU's energy savings. Overall, LithOS increases GPU efficiency, establishing a foundation for future OS research on GPUs.
This 12-Year-Old Sci-Fi Film Eerily Predicted Life in 2025. We Can Still Learn a Lot From It Today.
Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. I was 21 when I first watched Spike Jonze's 2013 sci-fi romance Her in theaters in New York City--a then–fresh college graduate teeming with the potent and deluded optimism that came with being a very broke and online millennial hoping to change the world. Her sparked some of my first reflections about whether tech innovation is inherently good or bad for society, and helped validate my early moral quandaries and panic at the time. I was graduating at the first turn of a recovering recession (mainly due to big tech investments in digital and social media) and securing my first full-time role as an online reporter. Though I was eager and rosy, a quiet, worried voice also began growing inside of me. Me, my job, my realities, were entirely dependent on tech--mainly Facebook content dissemination and programmatic turnkey digital ads--and I was not sure these huge tech investments by our broligarchical founding fathers would lead us anywhere good.